Search CORE

126 research outputs found

CAMISIM: Simulating metagenomes and microbial communities

Author: Belmann P
Bremges A
Dahms E
Darling AE
Demaere MZ
Dröge J
Fiedler J
Fritz A
Hofmann P
Lesker TR
Majda S
McHardy AC
Sczyrba A
Publication venue: 'Springer Science and Business Media LLC'
Publication date: 13/04/2018
Field of study

© 2019 The Author(s). Background: Shotgun metagenome data sets of microbial communities are highly diverse, not only due to the natural variation of the underlying biological systems, but also due to differences in laboratory protocols, replicate numbers, and sequencing technologies. Accordingly, to effectively assess the performance of metagenomic analysis software, a wide range of benchmark data sets are required. Results: We describe the CAMISIM microbial community and metagenome simulator. The software can model different microbial abundance profiles, multi-sample time series, and differential abundance studies, includes real and simulated strain-level diversity, and generates second- and third-generation sequencing data from taxonomic profiles or de novo. Gold standards are created for sequence assembly, genome binning, taxonomic binning, and taxonomic profiling. CAMSIM generated the benchmark data sets of the first CAMI challenge. For two simulated multi-sample data sets of the human and mouse gut microbiomes, we observed high functional congruence to the real data. As further applications, we investigated the effect of varying evolutionary genome divergence, sequencing depth, and read error profiles on two popular metagenome assemblers, MEGAHIT, and metaSPAdes, on several thousand small data sets generated with CAMISIM. Conclusions: CAMISIM can simulate a wide variety of microbial communities and metagenome data sets together with standards of truth for method evaluation

Helmholtz Zentrum für Infektionsforschung Repository

OPUS - University of Technology Sydney

Directory of Open Access Journals

Publications at Bielefeld University

Species-level functional profiling of metagenomes and metatranscriptomes.

Author: A Sczyrba
A Shafquat
AE Duran-Pinedo
AK Sharma
B Buchfink
B Langmead
BE Suzek
BK Swan
C Burke
C Luo
Curtis Huttenhower
D Medini
DH Huson
DT Truong
DT Truong
E Pasolli
EA Franzosa
EA Franzosa
Eric A. Franzosa
George Weingart
GG Silva
Gholamali Rahnavard
H Hauswedell
J Kim
J Lloyd-Price
J Lloyd-Price
J Ravel
J. Gregory Caporaso
JA Fuhrman
K Huang
Karen Schwarzberg Lipson
Lauren J. McIver
LR Thompson
LR Thompson
Luke R. Thompson
M Hamady
M Kanehisa
M Scholz
Melanie Schirmer
MY Galperin
N Segata
N Segata
Nicola Segata
OU Mason
P Petrenko
PJ Turnbaugh
R Caspi
RC Edgar
RD Finn
Rob Knight
S Abubucker
S Nayfach
S Sunagawa
S Sunagawa
T Bose
UniProt Consortium.
W Huang
Y Ye
Y Zhao
Publication venue: eScholarship, University of California
Publication date: 01/11/2018
Field of study

Functional profiles of microbial communities are typically generated using comprehensive metagenomic or metatranscriptomic sequence read searches, which are time-consuming, prone to spurious mapping, and often limited to community-level quantification. We developed HUMAnN2, a tiered search strategy that enables fast, accurate, and species-resolved functional profiling of host-associated and environmental communities. HUMAnN2 identifies a community's known species, aligns reads to their pangenomes, performs translated search on unclassified reads, and finally quantifies gene families and pathways. Relative to pure translated search, HUMAnN2 is faster and produces more accurate gene family profiles. We applied HUMAnN2 to study clinal variation in marine metabolism, ecological contribution patterns among human microbiome pathways, variation in species' genomic versus transcriptional contributions, and strain profiling. Further, we introduce 'contributional diversity' to explain patterns of ecological assembly across different microbial community types

Crossref

eScholarship - University of California

Decontamination of MDA Reagents for Single Cell Whole Genome Amplification

Author: A Raghunathan
AJ Varghese
Alexander Sczyrba
Christian Rinke
CY Ou
Damon Tighe
DE Brash
DR Zerbino
H Li
J Cadet
J Cadet
Jan-Fang Cheng
Janey Lee
K Zhang
Olivier Lespinet
PC Blainey
R Stepanauskas
Ramunas Stepanauskas
Rex Malmstrom
S Champlot
S Rodrigue
Scott Clingenpeel
T Woyke
T Woyke
Tanja Woyke
Y Marcy
Publication venue: Public Library of Science
Publication date: 01/01/2011
Field of study

Single cell genomics is a powerful and increasingly popular tool for studying the genetic make-up of uncultured microbes. A key challenge for successful single cell sequencing and analysis is the removal of exogenous DNA from whole genome amplification reagents. We found that UV irradiation of the multiple displacement amplification (MDA) reagents, including the Phi29 polymerase and random hexamer primers, effectively eliminates the amplification of contaminating DNA. The methodology is quick, simple, and highly effective, thus significantly improving whole genome amplification from single cells

Public Library of Science (PLOS)

CiteSeerX

Crossref

Directory of Open Access Journals

PubMed Central

Publications at Bielefeld University

Streaming histogram sketching for rapid microbiome analytics

Author: A Sczyrba
AG Shaw
AL Greninger
AP Carrieri
B Grüning
BD Ondov
C Alcon-Giner
C Kakkanatt
D Yang
DB Rusch
F Pedregosa
G Benoit
G Cormode
H Mulcahy-O’Grady
Human Microbiome Project Consortium
I Koychev
JD Forbes
K Sim
LP Coelho
LR Thompson
M Bawa
MW Libbrecht
Q Zhang
R Bovee
S Ioffe
S Seth
SY Anvar
T Brown
T Haveliwala
VB Dubinkina
W Wu
XC Morgan
Publication venue: 'Springer Science and Business Media LLC'
Publication date: 01/03/2019
Field of study

Background: The growth in publically available microbiome data in recent years has yielded an invaluable resource for genomic research, allowing for the design of new studies, augmentation of novel datasets and reanalysis of published works. This vast amount of microbiome data, as well as the widespread proliferation of microbiome research and the looming era of clinical metagenomics, means there is an urgent need to develop analytics that can process huge amounts of data in a short amount of time. To address this need, we propose a new method for the compact representation of microbiome sequencing data using similarity-preserving sketches of streaming k-mer spectra. These sketches allow for dissimilarity estimation, rapid microbiome catalogue searching and classification of microbiome samples in near real time. Results: We apply streaming histogram sketching to microbiome samples as a form of dimensionality reduction, creating a compressed ‘histosketch’ that can efficiently represent microbiome k-mer spectra. Using public microbiome datasets, we show that histosketches can be clustered by sample type using the pairwise Jaccard similarity estimation, consequently allowing for rapid microbiome similarity searches via a locality sensitive hashing indexing scheme. Furthermore, we use a ‘real life’ example to show that histosketches can train machine learning classifiers to accurately label microbiome samples. Specifically, using a collection of 108 novel microbiome samples from a cohort of premature neonates, we trained and tested a random forest classifier that could accurately predict whether the neonate had received antibiotic treatment (97% accuracy, 96% precision) and could subsequently be used to classify microbiome data streams in less than 3 s. Conclusions: Our method offers a new approach to rapidly process microbiome data streams, allowing samples to be rapidly clustered, indexed and classified. We also provide our implementation, Histosketching Using Little K-mers (HULK), which can histosketch a typical 2 GB microbiome in 50 s on a standard laptop using four cores, with the sketch occupying 3000 bytes of disk space

University of Liverpool Repository

Crossref

University of Birmingham Research Portal

Directory of Open Access Journals

Spiral - Imperial College Digital Repository

University of East Anglia digital repository

Recommended from our members

Expansion of the Genomic Encyclopedia of Bacteria and Archaea

Author: Cheng Jan-Fang
Eisen Jonathan A.
Hallam Steven
Hedlund Brian P.
Hugenholtz Philip
Inskeep William P.
Lee Janey
Liu Wen-Tso
Malfatti Stephanie
Rinke Christian
Sczyrba Alex
Sievert Stefan M.
Stepanauskas Ramunas
Tsiamis George
Woyke Tanja
Publication venue: Lawrence Berkeley National Laboratory
Publication date: 20/03/2011
Field of study

To date the vast majority of bacterial and archaeal genomes sequenced are of rather limited phylogenetic diversity as they were chosen based on their physiology and/ or medical importance. The Genomic Encyclopedia of Bacteria and Archaea (GEBA) project (Wu et al. 2009) is aimed at systematically filling the gaps of the tree of life with phylogenetically diverse reference genomes. However more than 99 percent of microorganisms elude current culturing attempts, severely limiting the ability to recover complete or even partial genomes of these largely mysterious species. These limitations gave rise to the GEBA uncultured project. Here we propose to use single cell genomics to massively expand the Genomic Encyclopedia of Bacteria and Archaea by targeting 80 single cell representatives of uncultured candidate phyla which have no or very few cultured representatives. Generating these reference genomes of uncultured microbes will dramatically increase the discovery rate of novel protein families and biological functions, shed light on the numerous underrepresented phyla that likely play important roles in the environment, and will assist in improving the reconstruction of the evolutionary history of Bacteria and Archaea. Moreover, these data will improve our ability to interpret metagenomics sequence data from diverse environments, which will be of tremendous value for microbial ecology and evolutionary studies to come

UNT Digital Library

Exploring neighborhoods in large metagenome assembly graphs reveals hidden sequence diversity

Author: A Satyanarayan
A Sczyrba
A Westbrook
BJ Tully
CT Brown
CW Beitel
D Li
D Standage
DH Parks
DH Parks
E Pasolli
ED Demaine
EI Olekhnovich
H Lin
I Sharon
J Chlebíková
J Koster
J Pell
J Towns
JD Hunter
K Katoh
M Kanehisa
RD Finn
RR Wick
S Nayfach
S Nurk
T Kluyver
TO Delmont
TP Barnum
Walt van der
Y Yang
Publication venue: 'Springer Science and Business Media LLC'
Publication date: 06/07/2020
Field of study

Genomes computationally inferred from large metagenomic data sets are often incomplete and may be missing functionally important content and strain variation. We introduce an information retrieval system for large metagenomic data sets that exploits the sparsity of DNA assembly graphs to efficiently extract subgraphs surround- ing an inferred genome. We apply this system to recover missing content from genome bins and show that substantial genomic se- quence variation is present in a real metagenome. Our software implementation is available at https://github.com/spacegraphcats/ spacegraphcats under the 3-Clause BSD License

Crossref

Birkbeck Institutional Research Online

A genomic catalog of Earth’s microbiomes

Author: Abreu H.
Acinas S. G.
Allen E.
Allen M. A.
Andersen G.
Anesio A. M.
Arkin A. P.
Attwood G.
Avila-Magaña V.
Badis Y.
Bailey J.
Baker B.
Baldrian P.
Barton H. A.
Beck D. A. C.
Becraft E. D.
Beller H. R.
Beman J. M.
Bernier-Latmani R.
Berry T. D.
Bertagnolli A.
Bertilsson S.
Bhatnagar J. M.
Bird J. T.
Blumer-Schuette S. E.
Bohannan B.
Borton M. A.
Brady A.
Brawley S. H.
Brodie J.
Brown S.
Brum J. R.
Brune A.
Bryant D. A.
Buchan A.
Buckley D. H.
Buongiorno J.
Cadillo-Quiroz H.
Caffrey S. M.
Campbell A. N.
Campbell B.
Carr S.
Carroll J. L.
Cary S. C.
Cates A. M.
Cattolico R. A.
Cavicchioli R.
Chen I.-M.
Chistoserdova L.
Chivian D.
Coleman M. L.
Consortium IMG/M Data
Constant P.
Conway J. M.
Correa de Souza R. S.
Crowe S.
Crump B.
Currie C.
Daly R.
Dehal P.
Denef V.
Denman S. E.
Desta A.
Dionisi H.
Dodsworth J.
Dombrowski N.
Donohue T.
Dopson M.
Driscoll T.
Dunfield P.
Dupont C. L.
Dynarski K. A.
Edgcomb V.
Edirisinghe J. N.
Edwards E. A.
Eloe-Fadrosh E. A.
Elshahed M. S.
Faria J. P.
Figueroa I.
Flood B.
Fortney N.
Fortunato C. S.
Francis C.
Gachon C. M. M.
Garcia S. L.
Gazitua M. C.
Gentry T.
Gerwick L.
Gharechahi J.
Girguis P.
Gladden J.
Gradoville M.
Grasby S. E.
Gravuer K.
Grettenberger C. L.
Gruninger R. J.
Guo J.
Habteselassie M. Y.
Hallam S. J.
Hatzenpichler R.
Hausmann B.
Hazen T. C.
Hedlund B.
Henny C.
Henry C. S.
Herfort L.
Hernandez M.
Hershey O. S.
Hess M.
Hollister E. B.
Hug L. A.
Hunt D.
Huntemann M.
Ivanova N. N.
Jansson J.
Jarett J.
Jungbluth S. P.
Kadnikov V. V.
Kelly C.
Kelly R.
Kelly W.
Kerfeld C. A.
Kimbrel J.
Kirton E.
Klassen J. L.
Konstantinidis K. T.
Kyrpides N. C.
Ladau J.
Lee L. L.
Li W.-J.
Loder A. J.
Loy A.
Lozada M.
Mac Cormack W. P.
MacGregor B.
Magnabosco C.
Maia de Oliveira V.
Maria da Silva A.
McKay R. M.
McMahon K.
McSweeney C. S.
Medina M.
Meredith L.
Mizzi J.
Mock T.
Momper L.
Moran M. A.
Morgan-Lang C.
Moser D.
Mouncey N. J.
Mukherjee S.
Muyzer G.
Myrold D.
Nash M.
Nayfach S.
Nesbø C. L.
Neumann A. P.
Neumann R. B.
Nielsen T.
Noguera D.
Northen T.
Norton J.
Nowinski B.
Nüsslein K.
Oliveira R. S.
Onstott T.
Osvatic J.
Ouyang Y.
O’Malley M. A.
Pachiadaki M.
Paez-Espino D.
Palaniappan K.
Parnell J.
Partida-Martinez L. P.
Peay K. G.
Pelletier D.
Peng X.
Pester M.
Pett-Ridge J.
Peura S.
Pjevac P.
Plominsky A. M.
Poehlein A.
Pope P. B.
Ravin N.
Reddy T. B. K.
Redmond M. C.
Reiss R.
Rich V.
Rinke C.
Rodrigues J. L. M.
Rossmassler K.
Roux S.
Sackett J.
Salekdeh G. H.
Saleska S.
Scarborough M.
Schachtman D.
Schadt C. W.
Schrenk M.
Schulz F.
Sczyrba A.
Sengupta A.
Seshadri R.
Setubal J. C.
Shade A.
Sharp C.
Sherman D. H.
Shubenkova O. V.
Sierra-Garcia I. N.
Simister R.
Simon H.
Sjöling S.
Slonczewski J.
Spear J. R.
Stegen J. C.
Stepanauskas R.
Stewart F.
Suen G.
Sullivan M.
Sumner D.
Swan B. K.
Swingley W.
Tarn J.
Taylor G. T.
Teeling H.
Tekere M.
Teske A.
Thomas T.
Thrash C.
Tiedje J.
Ting C. S.
Tringe S. G.
Tully B.
Tyson G.
Udwary D.
Ulloa O.
Valentine D. L.
Van Goethem M. W.
VanderGheynst J.
Varghese N.
Verbeke T. J.
Visel A.
Vollmers J.
Vuillemin A.
Waldo N. B.
Walsh D. A.
Weimer B. C.
Whitman T.
Wielen P. van der
Wilkins M.
Williams T. J.
Wood-Charlson E. M.
Woodcroft B.
Woolet J.
Woyke T.
Wrighton K.
Wu D.
Ye J.
Young E. B.
Youssef N. H.
Yu F. B.
Zemskaya T. I.
Ziels R.
Publication venue: Nature Research
Publication date: 23/11/2020
Field of study

The reconstruction of bacterial and archaeal genomes from shotgun metagenomes has enabled insights into the ecology and evolution of environmental and host-associated microbiomes. Here we applied this approach to >10,000 metagenomes collected from diverse habitats covering all of Earth’s continents and oceans, including metagenomes from human and animal hosts, engineered environments, and natural and agricultural soils, to capture extant microbial, metabolic and functional potential. This comprehensive catalog includes 52,515 metagenome-assembled genomes representing 12,556 novel candidate species-level operational taxonomic units spanning 135 phyla. The catalog expands the known phylogenetic diversity of bacteria and archaea by 44% and is broadly available for streamlined comparative analyses, interactive exploration, metabolic modeling and bulk download. We demonstrate the utility of this collection for understanding secondary-metabolite biosynthetic potential and for resolving thousands of new host linkages to uncultivated viruses. This resource underscores the value of genome-centric approaches for revealing genomic properties of uncultivated microorganisms that affect ecosystem processes

KITopen

ViennaRNA Package 2.0

Author: A Busch
A Sczyrba
A Waugh
AJ Enright
AR Gruber
AR Gruber
AR Gruber
B Kaczkowski
B Knudsen
B Matthews
C Aksay
C Flamm
C Flamm
C Flamm
C Höner zu Siederdissen
CB Do
Christian Höner zu Siederdissen
Christoph Flamm
D Sankoff
D Thirumalai
D Upper
DA Benson
DH Mathews
DH Mathews
DH Mathews
DH Turner
EP Nawrocki
H Kiryu
H Tafer
H Tafer
H Tafer
H Tafer
Hakim Tafer
I Tinoco Jr
I Tinoco Jr
IL Hofacker
IL Hofacker
IL Hofacker
IL Hofacker
IL Hofacker
IL Hofacker
IL Hofacker
IL Hofacker
IL Hofacker
Ivo L Hofacker
J Hertel
J Hertel
J Reeder
J SantaLucia
JA Jaeger
JH Havgaard
JN Zadeh
JN Zadeh
JS McCaskill
JS Reuter
K Darty
K Reiche
L Dagum
L He
M Andronescu
M Andronescu
M Andronescu
M Fekete
M Hamada
M Höchsmann
M Kalaš
M Larkin
M Parisien
M Rehmsmeier
M Tacker
M Zuker
M Zuker
M Zuker
MS Andronescu
MS Waterman
NR Markham
P Gardner
P Gardner
P Schuster
Peter F Stadler
R Dowell
R Klein
R Lorenz
R Nussinov
R Nussinov
R Thadani
RA Dimitrov
RE Bruccoleri
Ronny Lorenz
RR Stocsits
S Bernhart
S Bonhoeffer
S Heyne
S Washietl
S Will
S Wuchty
S Zakov
SH Bernhart
SH Bernhart
SM Freier
SR Eddy
Stephan H Bernhart
T Xia
U Mückstein
V Rusinov
W Beyer
W Fontana
W Fontana
W Fontana
W Fontana
W Pearson
Y Ding
Y Ding
Publication venue: BioMed Central
Publication date: 01/01/2011
Field of study

Abstract Background Secondary structure forms an important intermediate level of description of nucleic acids that encapsulates the dominating part of the folding energy, is often well conserved in evolution, and is routinely used as a basis to explain experimental findings. Based on carefully measured thermodynamic parameters, exact dynamic programming algorithms can be used to compute ground states, base pairing probabilities, as well as thermodynamic properties. Results The <monospace>ViennaRNA</monospace> Package has been a widely used compilation of RNA secondary structure related computer programs for nearly two decades. Major changes in the structure of the standard energy model, the <it>Turner 2004 </it>parameters, the pervasive use of multi-core CPUs, and an increasing number of algorithmic variants prompted a major technical overhaul of both the underlying <monospace>RNAlib</monospace> and the interactive user programs. New features include an expanded repertoire of tools to assess RNA-RNA interactions and restricted ensembles of structures, additional output information such as <it>centroid </it>structures and <it>maximum expected accuracy </it>structures derived from base pairing probabilities, or <it>z</it>-<it>scores </it>for locally stable secondary structures, and support for input in <monospace>fasta</monospace> format. Updates were implemented without compromising the computational efficiency of the core algorithms and ensuring compatibility with earlier versions. Conclusions The <monospace>ViennaRNA Package 2.0</monospace>, supporting concurrent computations <monospace>via OpenMP</monospace>, can be downloaded from <url>http://www.tbi.univie.ac.at/RNA</url>.</p

Crossref

Springer - Publisher Connector

Directory of Open Access Journals

Fraunhofer-ePrints

PubMed Central

Permanent Hosting, Archiving and Indexing of Digital Resources and Assets

Critical Assessment of Metagenome Interpretation:A benchmark of metagenomics software

Author: A Mikheenko
Aaron E Darling
Adrian Fritz
Alexander Sczyrba
Alexey Gurevich
Alice C McHardy
Andreas Bremges
B Liu
Bernhard Y Renard
Bertrand Denis
Burton K H Chia
C Lozupone
Charles Deltel
Chirag Jain
Christopher Quince
Claire Lemaitre
D Coil
D Koslicki
D Koslicki
D Koslicki
D Li
D Turaev
Daniel A Cuevas
David Koslicki
DD Kang
DE Wood
DH Huson
Dmitrij Turaev
Dominique Lavenier
Dongwan Don Kang
E Pruesse
Edward M Rubin
Eik Dahms
Fernando Meyer
Genivaldo Gueiros Z Silva
GG Silva
Guillaume Rizk
H Klingenberg
Hans-Peter Klenk
Heiner Klingenberg
HH Lin
Hsin-Hung Lin
I Gregor
Ivan Gregor
J Alneberg
J Dröge
JA Chapman
Jeff L Froula
Jeffrey J Cook
Jessika Fiedler
Johannes Dröge
Julia A Vorholt
K Mavromatis
KT Konstantinidis
Lars Hestbjerg Hansen
M Arumugam
M Balvočiūtė
M Strous
M Yassour
Marc Strous
Markus Göker
Matthew Z DeMaere
Michael Beckstette
Michael D Barton
Mihai Pop
ML Bendall
Monika Balvočiūtė
N Kashtan
N Sangwan
N Segata
Nicole Shapiro
Nikos C Kyrpides
Niranjan Nagarajan
NP Nguyen
O Koren
P Belmann
Paul Schulze-Lefert
Peter Belmann
Peter Hofmann
Peter Meinicke
Philip D Blood
Pierre Peterlongo
R Chikhi
R Ounit
Rayan Chikhi
Robert A Edwards
Robert Egan
RR Miller
Ruben Garrido-Oter
S Boisvert
S Chatterjee
S Gao
S Lindgreen
S Sunagawa
Stefan Janssen
Stephan Majda
Steven W Singer
Surya Saha
Søren J Sørensen
T Thomas
Tanja Woyke
Thomas Lingner
Thomas Rattei
Tue Sparholt Jørgensen
V Marx
VC Piro
Vitor C Piro
Y Bai
Yang Bai
Yu-Chieh Liao
Yu-Wei Wu
YW Wu
Zhong Wang
Publication venue: 'Springer Science and Business Media LLC'
Publication date: 01/01/2017
Field of study

International audienceIn metagenome analysis, computational methods for assembly, taxonomic profilingand binning are key components facilitating downstream biological datainterpretation. However, a lack of consensus about benchmarking datasets andevaluation metrics complicates proper performance assessment. The CriticalAssessment of Metagenome Interpretation (CAMI) challenge has engaged the globaldeveloper community to benchmark their programs on datasets of unprecedentedcomplexity and realism. Benchmark metagenomes were generated from newlysequenced ~700 microorganisms and ~600 novel viruses and plasmids, includinggenomes with varying degrees of relatedness to each other and to publicly availableones and representing common experimental setups. Across all datasets, assemblyand genome binning programs performed well for species represented by individualgenomes, while performance was substantially affected by the presence of relatedstrains. Taxonomic profiling and binning programs were proficient at high taxonomicranks, with a notable performance decrease below the family level. Parametersettings substantially impacted performances, underscoring the importance ofprogram reproducibility. While highlighting current challenges in computationalmetagenomics, the CAMI results provide a roadmap for software selection to answerspecific research questions

Roskilde Universitet

HAL Descartes

Warwick Research Archives Portal Repository

MPG.PuRe

Hal-Diderot

Repository for Publications and Research Data

Crossref

National Health Research Institues

OPUS - University of Technology Sydney

INRIA a CCSD electronic archive server

Copenhagen University Research Information System

eScholarship - University of California

Publications at Bielefeld University

University of East Anglia digital repository

ScholarBank@NUS

HAL-Rennes 1

Large expert-curated database for benchmarking document similarity detection in biomedical literature search

Author: Aanei CM
Abid MB
Abramowitz MK
Abu-Zaid A
Afnan M
Agarabi C
Ahmad R
Aizat WM
Al-Farha AA
Al-Lawama M
Alanio A
Alaux C
Albiol J
Albrecht DR
Albuquerque LG
Alimba CG
Allardyce J
Almeida GMF
Alonso-Caneiro D
Alper OM
Amer SEDR
Amiya E
Ammerman BA
Amorim RM
An Q
Andersen SU
Aplin JD
Argyropoulos C
Armitage C
Ascher DB
Ashry M
Asmann YW
Assaeed AM
Atack JM
Atanasov AG
Atchison DA
Atkins GJ
Atlas L
Avery SV
Avillach P
Baade PD
Backman L
Badie C
Bae T
Baier D
Baker CI
Bakkach J
Baldi A
Ball E
Bannon R
Bansal A
Bardot O
Barnett AG
Barraud P
Basharat Z
Basner M
Batra J
Baumert P
Bazanova OM
Beale A
Beck CR
Becker D
Beddoe T
Bell ML
Benezeth Y
Bengtsson-Palme J
Berbesque C
Berezikov E
Bergsland N
Berners-Price S
Bernhardt P
Berrevoet F
Berry E
Berthold M
Bessa TB
Beyene TJ
Biedermann PHW
Bijleveld E
Billington C
Birch J
Bittner F
Bitzer M
Blakely RD
Blanck O
Blaskovich MAT
Bleackley M
Blombach F
Blum R
Boehme KA
Boelaert M
Bogdanos D
Bonvin AMJJ
Bosch C
Bosch O
Boudreau SA
Bourgoin T
Bourke E
Bouvard D
Boykin LM
Bradley G
Bradshaw W
Bramoweth AD
Brand T
Braubach O
Braun D
Braun RJ
Brenneisen P
Bridges KM
Brown JAL
Brown P
Browngardt C
Brownlie J
Bruhl A
Bukowy-Bieryllo Z
Bull JA
Burt A
Bush SJ
Butler LM
Byrareddy SN
Byrne HJ
Cabantous S
Cai Y
Calatayud S
Campana LG
Campbell M
Candal E
Cao Z
Cao Z
Cardoso P
Carlson K
Carter D
Cascella M
Casillas S
Castelvetro V
Caswell PT
Catry T
Cavalli G
Cernava T
Cerovsky V
Chacko G
Chagoyen M
Chakraborty S
Chan SS
Chandrasekaran AR
Chatzitheochari S
Chavez-Fumagalli MA
Chen B
Chen C-E
Chen C-S
Chen DF
Chen H
Chen H
Chen J-T
Chen X
Chen Y
Cheng C
Cheng J
Cheng S
Cheung JTK
Chinapaw M
Chinopoulos C
Cho WCS
Chong L
Chowdhury D
Chung H-J
Chwalibog A
Ciresi A
Cobine PA
Cockcroft S
Coelho LP
Colella V
Conesa A
Conway A
Cook PA
Cooper DN
Cooper J
Coqueret O
Corea EM
Cosacak MI
Costa BM
Costa E
Costa VD
Coupland C
Crawford SY
Cruz AD
Cui H
Cui Q
Cuiv PO
Culver DC
Cuypers M
Cyr N
D'Angiulli A
Dahms TES
Dai Z
Daigle F
Dalgleish R
Dalrymple BP
Danchin A
Danielsen HE
Darras S
Daulatabad SV
Davidson SM
Day DA
de Keersmaecker K
de Leeuw F-E
Dean LT
Debrabant B
Degirmenci V
del Tredici AL
Delahay RM
Demaison L
Denzel MS
Deschodt M
Devkota HP
Devriendt K
Dhariwal R
Diao J
Ding J
Dings RPM
Diouf B
Dixon R
Dlamini SV
Dogan Y
Domingues HS
Dong XC
Donner CF
Dono M
Doxey AC
Dressick W
Drevon CA
Duan H
Ducho C
Ducommun B
Dudley KJ
Dufies M
Duijf PHG
Dumaz N
Dwarakanath BS
Ebell MH
Echeverria N
Ecke T
Eckweiler D
Eerola T
Effiong A
Ehret F
Eisenhardt S
Eixarch E
El-Adawy H
El-Esawi MA
Elkum N
Emmrich JV
Engel MS
Engel N
Epp T
Erickson TB
Esfahlani SS
Eskelinen E-L
Eskew EA
Esnakul AK
Eustace AJ
Evangelou E
Fairhead M
Falk S
Fallah M
Falter-Wagner CM
Fan X
Farber DB
Faville MJ
Feghali KA
Fejzo MS
Fernandez-Triana J
Festa F
Feteira A
Feyerabend F
Fierz W
Filipp FV
Fiona .
Flegel WA
Flood-Page P
Florio T
Forano E
Forsayeth J
Fox SA
Franks SJ
Frentiu FD
Friebe M
Frilander MJ
Fu X
Fujita S
Furuta S
Fuss J
Gabrielsen M
Gajda M
Galea I
Galluzzi L
Gani F
Ganpule AP
Gao J
Garcia-Alix A
Gatchell M
Gaullier G
Gedye K
Gelfer Y
Ghelardi E
Gill MR
Gilliham M
Giordano M
Giunta C
Gladue DP
Gleeson PA
Gloyn L
Gnasso A
Goarant C
Gobet A
Goggs R
Gong H
Gonzalezlez-Prendes R
Goodin A
Goodyear CS
Gora D
Gough MJ
Govender P
Govinden U
Goyal R
Graham EB
Graham KE
Grande-Perez A
Graves PM
Greene G
Greenwald NF
Greidanus H
Greiff V
Grice D
Grimm DG
Groen EJN
Gruber J
Grunau C
Grundle DS
Gruneberg P
Grybos M
Guisado JL
Gumede N
Gumulya Y
Guo Y
Gurevich VV
Gurney-Champion OJ
Gusev O
Gutierrez-Sacristan A
Habes M
Hacker E
Hage SR
Hagen G
Hahn S
Haller DM
Hammerschmidt S
Han H
Han J
Han Q
Han R
Handfield M
Hanson J
Haore G
Hapuarachchi HC
Harder T
Hardingham JE
Harrison P
Hartmann MD
Harvey DJ
Haston S
Heck M
Heers M
Heffler E
Heinrich M
Helantera H
Herbelet S
Hew KF
Higginbottom DB
Higuchi Y
Hilton R
Hiroi N
Hobbs E
Hodzic E
Hoenner X
Hojsgaard D
Hone A
Hongoh Y
Honjo K
Horbar J
Hori H
Hu G
Hu P
Huber HP
Huber M
Hueso LE
Huirne J
Hurt L
Huttner FJ
Idborg H
Ide K
Ikeo K
Ikonomopoulou MP
Ingley E
Jakeman PM
Janga SC
Janzen T
Jayaraman J
Jeltsch A
Jensen A
Jeurissen P
Jia H
Jia H
Jia S
Jiang F
Jiang J
Jiang X
Jibb LA
Jin Y
Jo D
Johnson AM
Johnson DM
Johnston M
Jongen S
Jonscher KR
Jorens PG
Jorgensen JOL
Josse C
Joubert JW
Jung S-H
Junior AM
Jurman G
Kabra D
Kahan T
Kaiser S
Kamagata K
Kamboj SK
Kamiya H
Kane NC
Kang Y-K
Karamanos Y
Karmakar C
Karp NA
Kasian O
Kauppila JH
Kaye LK
Kelly R
Kelly S
Kenna R
Kennedy J
Kersten B
Khalaf RA
Khalid JM
Khan MM
Khatlani T
Khider T
Kijanka GS
Kim Y-M
King SRB
Kinyanjui T
Kish JK
Klempnauer K-H
Kleppe A
Klump H
Kluz T
Knox P
Kobayashi T
Kobold S
Koch K-W
Kohanbash G
Kohls G
Kohonen-Corish MRJ
Koleva-Kolarova RG
Kong X
Konkle-Parker D
Korpela KM
Kostrikis LG
Kraiczy P
Kratz H
Krause G
Krebsbach PH
Kristensen SR
Kristiansson E
Kueberuwa G
Kugler J-M
Kulkarni A
Kumar G
Kumar N
Kumar N
Kumari P
Kunimatsu A
Kurdak H
Kurgan L
Kurniawan NA
Kwon YD
Lachat C
Lacy-Colson J
Lagisz M
Lai HM
Laky B
Lalaouna D
Lammerding J
Lange M
Larrosa M
Laslett AL
Latif A
Lau CL
Lauschke VM
LeClair EE
Lee K-W
Lee M-S
Lee M-Y
Lee S
Li B
Li G
Li J
Li J
Li J
Li Z
Liang D
Liang S
Lidbury BA
Lieb K
Liehr T
Liew AWC
Lim CJ
Lim YY
Lin MZ
Lindsey ML
Line P-D
Liu D
Liu E
Liu F
Liu F
Liu H
Liu H
Liu S
Liu X
Liu Y-P
Lloyd VK
Lo T-W
Locci E
Loft ND
Loidl J
Lopez-Escamez JA
Lopez-Ruiz FJ
Lorenzen J
Lorkowski S
Lovell NH
Lu H
Lu J-J
Lu Q
Lu W
Lu Z
Luengo GS
Lund BA
Lundh L-G
Lussier AA
Luu AM
Lynch I
Lysy PA
Ma C
Ma L
Ma L
Ma L
Ma R
Ma W
Mabb A
Mack HG
Mackey DA
Mahavadi P
Mahdavi SR
Maher P
Maher T
Maibach EW
Maity SN
Malgrange B
Mamoulakis C
Mangoni AA
Manke T
Manstead ASR
Mantalaris A
Marchbank KJ
Marinello F
Marsal J
Marschalek R
Marschall H-U
Martin CS
Martin FL
Martinez-Raga J
Martinez-Salas E
Martis E
Marzocchi U
Mather DE
Mathieu D
Matsui Y
Maza E
McCrum C
McCutcheon JE
McGarrigle CA
Mckay GJ
McMillan B
McMillan N
Meads C
Medina L
Merrick BA
Meseko C
Metzger DW
Meule A
Meunier FA
Michaelis M
Micheau O
Miele AE
Mier P
Mihara H
Min R
Mintz EM
Miotla P
Mitchell KM
Mizukami T
Moal I
Moalic Y
Mohapatra DP
Molari M
Molleman L
Mondal SR
Montagutelli X
Monteiro A
Montes M
Moore MD
Moran JV
Morcillo E
Morozov SY
Mort M
Moss WN
Moultos OA
Moyer R
Mukherjee M
Murai N
Murphy DJ
Murphy SK
Murray SA
Muth T
Naganawa S
Nagler K
Nakayama K
Nammi S
Nandakumar KS
Narayan E
Nasios G
Natoli RM
Navaratnarajah .
Neumann P-A
Ng G
Nguyen F
Nicol C
Nicoletti R
Nie J
Nie Y
Niehof M
Niemeyer F
Nilsen EB
Nilsson H
Nixon B
Nobile CJ
Norris AD
Nwaiwu O
O'Mahony M
O'Toole R
Ogami K
Ohgami RS
Ohlsson S
Ohtomo T
Olatunbosun O
Oldenmenger WH
Olofsson P
Olumayede E
Orme MW
Ortiz A
Oster H
Ostrikov K
Otto S
Ou J
Outeiro TF
Ouyang S
Paganoni S
Page A
Pallebage-Gamarallage M
Palm C
Palma J-A
Pan Z
Panthee S
Paradies Y
Parchi P
Parsons JR
Parsons MH
Parsons N
Pascal P
Paterson R
Paul E
Pearce SP
Pearson JA
Peckham M
Pedemonte N
Peifer M
Pelkonen T
Pelleri MC
Pellizzon MA
Peng Y
Perco P
Pereira JL
Peres MA
Petrelli M
Pheko M
Pichugin A
Pinto CJC
Pinto IM
Pinto KA
Piotrowski M
Piovesan A
Plevris JN
Pluess M
Podolsky IM
Pollesello P
Polz M
Ponti G
Popoola SI
Porcelli P
Portilla M
Portillo MC
Pourret O
Prajapati AS
Pranata R
Prescott J
Prieto D
Prince M
Pritchard AL
Pusch S
Qi D
Qi X
Quinn GP
Quinn TJ
Raghava GPS
Rahimi F
Rahman MS
Raikou VD
Ramula S
Ranft A
Rappsilber J
Reddan T
Rehfeldt F
Reiling JH
Remacle C
Reschke CR
Rezaei M
Rhodes J
Riddick EW
Ritter U
Riva G
Roach NW
Roberts DD
Roberts NJ
Robles G
Rodrigues T
Rodriguez C
Roislien J
Roobol MJ
Ross K
Ross SA
Rotge J-Y
Rowe AD
Rowe JA
Ruepp A
Rust P
Saad S
Sabnis SC
Sack GH
Saggar M
Saito Y
Salama MF
Sallmon H
Santos M
Saudemont A
Sava G
Schrading S
Schramm A
Schreiber M
Schuele B
Schuler S
Schulte LN
Schuon RA
Schymkowitz J
Sczyrba A
Seib KL
Senghore T
Seow E
Sergeant K
Shabalin IG
Shahid S
Shalchyan V
Shen J
Shi H-P
Shimada T
Shin J-S
Shortt C
Siebers R
Sillanpaa E
Silveyra P
Skinner D
Small I
Smeets PAM
Smith SS
So P-W
Solano F
Sonenshine DE
Song H
Song J
Sorzano CO
Southall T
Speakman JR
Srinivasan MV
St Hilaire C
Stabile LP
Staege MS
Stasiak A
Steadman KJ
Stein N
Stella A
Stephens AW
Stevanovic D
Stewart CJ
Stewart DI
Stine K
Storlazzi C
Stoynova NV
Strzalka W
Suarez OM
Subhash S
Sukocheva O
Sultana T
Sumant AV
Summers MJ
Sun G
Sydes M
Tacon P
Tamaian R
Tan A-C
Tan E-C
Tan K-H
Tanaka K
Tang H
Tanino Y
Targett-Adams P
Tayebi M
Tayyem R
Tebbe CC
Telfer EE
Tempel W
Teodorczyk-Injeyan JA
Terrier O
Testoni I
Thijs G
Thorne S
Thrift AG
Tiffon C
Tinnefeld P
Tjahjono DH
Tofani M
Tolle F
Torga G
Toth E
Tressoldi P
Troder SE
Tsapas A
Tsirigotis K
Turak A
Tuttle N
Tzotzos G
Uchendu F
Udo EE
Uhle F
Utsumi T
Uversky VN
Vaidyanathan S
Vaillant M
Valsesia A
Van de Mortel T
Van den Bos W
van Meerten T
van Nieuwerburgh F
van Raaij MJ
van Ruitenbeek J
Vandenbroucke RE
Vanneste S
Veiga FH
Vendrell M
Verloh N
Vesk PA
Vickers P
Victor VM
Villemur R
Villet MH
Vindin H
Viveiros M
Vohl M-C
Voolstra CR
Vorholt JA
Voskarides K
Voutchkova DD
Vuillemin A
Wakelin S
Waldron L
Walsh LJ
Wang AY
Wang F
Wang Y
Watanabe Y
Weigert A
Weinstock C
Wen J-C
Werner GDA
Werten S
Westermair AL
Wham C
White EP
Widera D
Wiener J
Wilharm G
Wilkinson S
Williams R
Willmann R
Wilson C
Wirth B
Wojan TR
Woldesemayat AA
Wolff M
Wong A
Wong BM
Wu T-W
Wuerbel H
Xia W
Xiao X
Xu D
Xu H
Xu J
Xu J
Xu JW
Xue B
Xue Y
Yadollahpour A
Yalcin S
Yamato M
Yan H
Yang E-C
Yang H
Yang L
Yang S
Yang SY
Yang W
Yang Y
Ye Y
Ye Z-Q
Yeung AWK
Yin C-C
Yli-Kauhaluoma J
Yoneyama H
Yu Y
Yuan G-C
Yuh C-H
Zabetakis I
Zaccolo M
Zaucha J
Zeng C
Zeng E
Zevnik B
Zhang C
Zhang C
Zhang J
Zhang L
Zhang L
Zhang X
Zhang Y
Zhang Y
Zhang Z
Zhang Z
Zhang Z-Y
Zhao X
Zhao Y
Zhou K
Zhou M
Zhu S
Ziegler A
Zinke K
Zuberbier T
Publication venue: OXFORD UNIV PRESS
Publication date: 29/10/2019
Field of study

Document recommendation systems for locating relevant literature have mostly relied on methods developed a decade ago. This is largely due to the lack of a large offline gold-standard benchmark of relevant documents that cover a variety of research fields such that newly developed literature search techniques can be compared, improved and translated into practice. To overcome this bottleneck, we have established the RElevant LIterature SearcH consortium consisting of more than 1500 scientists from 84 countries, who have collectively annotated the relevance of over 180 000 PubMed-listed articles with regard to their respective seed (input) article/s. The majority of annotations were contributed by highly experienced, original authors of the seed articles. The collected data cover 76% of all unique PubMed Medical Subject Headings descriptors. No systematic biases were observed across different experience levels, research fields or time spent on annotations. More importantly, annotations of the same document pairs contributed by different scientists were highly concordant. We further show that the three representative baseline methods used to generate recommended articles for evaluation (Okapi Best Matching 25, Term Frequency–Inverse Document Frequency and PubMed Related Articles) had similar overall performances. Additionally, we found that these methods each tend to produce distinct collections of recommended articles, suggesting that a hybrid method may be required to completely capture all relevant articles. The established database server located at https://relishdb.ict.griffith.edu.au is freely available for the downloading of annotation data and the blind testing of new methods. We expect that this benchmark will be useful for stimulating the development of new powerful techniques for title and title/abstract-based search engines for relevant articles in biomedical research

UCL Discovery